Papers with German text
Muted: Multilingual Targeted Offensive Speech Identification and Visualization (2023.emnlp-demo)
Copied to clipboard
Christoph Tillmann, Aashka Trivedi, Sara Rosenthal, Santosh Borse, Rong Zhang, Avirup Sil, Bishwaranjan Bhattacharjee
| Challenge: | Existing visualizations of offensive language use only sentence level annotations, but there are few that explore spans and other languages. |
| Approach: | They propose a system to identify multilingual HAP content by displaying offensive arguments and their targets using heat maps to indicate their intensity. |
| Outcome: | The proposed model can identify toxic spans without further fine-tuning using existing models and its attention mechanism out-of-the-box. |
Subjective Text Complexity Assessment for German (2022.lrec-1)
Copied to clipboard
| Challenge: | Often, readability is defined as how easily a written text is to read. |
| Approach: | They propose to use a corpus of sentences provided by a German IT service provider to assess the readability of German text. |
| Outcome: | The proposed model can predict complexity of German text by using linguistically motivated features. |
Stylometry in a Bilingual Setup (2020.lrec-1)
Copied to clipboard
| Challenge: | a stylometric method of comparing texts by most frequent words does not allow direct comparison of original texts and their translations, i.e. across languages. |
| Approach: | They propose a stylometric method that removes language-specific features and parses each language counterpart with a corresponding language model in UDPipe. |
| Outcome: | The proposed method removes language-specific features and keeps linguistically independent features of individual author signal. |
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)
Copied to clipboard
Michel Plüss, Jan Deriu, Yanick Schraner, Claudio Paonessa, Julia Hartmann, Larissa Schmidt, Christian Scheller, Manuela Hürlimann, Tanja Samardžić, Manfred Vogel, Mark Cieliebak
| Challenge: | We present a corpus of Swiss German speech annotated with Standard German text at the sentence level. |
| Approach: | They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them . |
| Outcome: | The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date . |
DGS-Fabeln-1: A Multi-Angle Parallel Corpus of Fairy Tales between German Sign Language and German Text (2024.lrec-main)
Copied to clipboard
Fabrizio Nunnari, Eleftherios Avramidis, Cristina España-Bonet, Marco González, Anna Hennes, Patrick Gebhard
| Challenge: | a parallel corpus of German text and videos containing fairy tales interpreted into the German Sign Language (DGS) is the first corpus filmed from 7 angles and one of the few sign language corpora globally which have been filmed simultaneously. |
| Approach: | They present a parallel corpus of German fairy tales interpreted by a native DGS signer. |
| Outcome: | The proposed corpus is the first semi-naturally expressed DGS that has been filmed from 7 angles and where the listener has been simultaneously filmed. |
LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of sentence-aligned triples of German audio, German text, and English translation is available for speech recognition . a large corpus is available to date for end-to-end speech translation based on parallel data . |
| Approach: | They present a corpus of sentence-aligned triples of German audio, German text, and English translation based on German audio books. |
| Outcome: | The proposed corpus is the largest resource for German speech recognition and for end-to-end German-to English speech translation. |